Papers with fine-grained evaluation protocol

2 papers
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluations of retrieval-augmented generation systems are limited . sub-question coverage measures how well a RAG system addresses different facets of a question.
Approach: They propose a framework for evaluation based on sub-question coverage . they propose to decompose questions into sub-questions and classify them into three types .
Outcome: The proposed evaluation framework measures how well a RAG system addresses different facets of a question.
Benchmarking Deflection and Hallucination in Large Vision-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks overlook conflicts between visual and textual evidence and the importance of generating deflections when incomplete knowledge is retrieved.
Approach: They propose a dynamic curation pipeline that preserves benchmark difficulty over time . they propose 'vlm-DeflectionBench' benchmark to probe model behaviour under conflicting evidence .
Outcome: The proposed benchmarks overlook conflicts between visual and textual evidence and are prone to obsolescence . the proposed benchmark is based on 2,775 samples spanning diverse retrieval settings .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations